The Journal of Molecular Diagnostics
○ Elsevier BV
Preprints posted in the last 90 days, ranked by how well they match The Journal of Molecular Diagnostics's content profile, based on 39 papers previously published here. The average preprint has a 0.03% match score for this journal, so anything above that is already an above-average fit.
Feierabend, S.; Künstner, A.; Forster, M.; Helbing, T.; Gebauer, N.; Gemoll, T.; Axt, F.; Nimmagadda, S. C.; Ranganathan, L.; Schwandt, J.; Heber, M.; Szymczak, S.; Hohensee, I.; Fliedner, S. M. J.; Scherer, F.; Oberländer, M.; Derer-Petersen, S.; Busch, H.; von Bubnoff, N.; Dazert, E.
Show abstract
Cancer treatment has shifted toward personalized therapy based on molecular profiling, particularly in advanced disease. Existing circulating tumor DNA panels are often broad, generating many non-actionable variants and incurring costs that limit routine use in molecular tumor boards. We developed and validated a manufacturer-independent, 109-gene liquid biopsy-centered pan-cancer open next generation sequencing panel (LION panel), combined with an in-house bioinformatic pipeline to support clinical decision-making. A total of 87 samples were analyzed, including 17 reference samples, 21 healthy blood donor controls, and 49 patient samples including nine tumor entities. The LION panel achieved 92% sensitivity and 99% specificity in reference samples, with high concordance to digital droplet PCR (r = 0.99). It detected variant allele frequencies as low as 0.05% (tumor-informed) and 0.5% (tumor-uninformed). Clinical concordance reached 82% with blood-based digital droplet PCR and 75% with whole exome tissue sequencing. In representative cases, variant dynamics correlated with disease progression and revealed additional targetable variants. Overall, the LION panel supports clinical decision-making by enabling identification of targetable variants, disease monitoring, and detection of treatment resistance, particularly when tumor tissue is unavailable.
Marvin, C. T.; Devaney, J. M.; Buckingham, K. J.; Noya, J.; Shively, K. M.; Jacques, C.; Galey, M.; Storz, S. H.; Goffena, J.; Berlyoung, A. S.; Patterson, K. E.; Shaffer, T.; Zakarian, C.; McGee, S. R.; Smith, J. D.; Lochovsky, L.; Gustafson, J. A.; Sommerland, O. M.; Anderson, K.; Love-Nichols, J.; Facio, F. M.; Robertson, A. V.; Rowell, W. J.; Lake, J. A.; Carroll, A.; Miller, D. E.; Wei, C. L.; McWalter, K.; Wenger, T. L.; University of Washington Center for Rare Disease Research, ; Johnson, B.; Bamshad, M. J.; Chong, J. X.
Show abstract
Long-read whole genome sequencing (lrWGS) shows promise as an all-in-one test to detect clinically relevant variants and variants difficult to detect by current short-read whole genome sequencing (srWGS) pipelines. Comparisons between lrWGS and srWGS (or exome sequencing) pipelines will become commonplace as lrWGS is more widely adopted for clinical testing, particularly for individuals not diagnosed by srWGS. However, the sensitivity of lrWGS for detecting variants previously identified and prioritized by clinical srWGS has yet to be assessed. As part of the SeqFirst-neo study, a subset of critically ill newborns and their parents who underwent clinical srWGS also underwent lrWGS on the Oxford Nanopore Technologies (ONT) and Pacific Biosciences (PacBio) platforms. In total, 134 families were sequenced across multiple technologies including 128 families with clinical srWGS who were sequenced on both lrWGS platforms. We compared the variants reported by clinical testing with the variants identified by lrWGS. Among the 128 families sequenced on all three platforms, 89 SNV/indels and 14 SV/CNVs clinically reported by the srWGS testing pipeline were evaluated. All variants assessed in probands were ultimately detected by both lrWGS platforms, although three events were not detected prior to application of an updated variant caller, highlighting the rapid evolution of lrWGS variant calling. Additionally, breakpoint coordinates and event sizes often differed substantially between calls from srWGS and events called in lrWGS data. Our work demonstrates that while most clinically reported variants from srWGS can be detected by lrWGS pipelines, challenges remain when attempting direct comparisons, particularly for SV/CNVs.
Burssed, B.; van der Sanden, B.; Hops, W.; Neveling, K.; Kamping, E.; van Beek, R.; den Ouden, A.; Derks, R.; Timmermans, R.; Perrone, E.; Ramos, M. A.; Bellucco, F. T.; Hoischen, A.; Melaragno, M. I.
Show abstract
Complex rearrangements are one of the rarest types of structural variants (SVs) and can be divided into two categories: complex chromosomal rearrangements (CCRs) and complex genomic rearrangements (CGRs). CCRs include structural rearrangements that present at least three breakpoints and show exchange of genetic material between more than two chromosomes and CGRs are rearrangements that present more than one junction and/or more than one SV in cis. They are usually formed by one of the chromoanagenesis mechanisms, where a massive disruptive cellular event leads to multiple structural rearrangements. Classical cytogenomic techniques have been commonly applied for their characterization, but methodologies that involve longer DNA molecules, namely optical genome mapping (OGM) and long-read genome sequencing (lrGS), present a considerably higher SV detection resolution, revealing more details about the rearrangements, including precise breakpoint location. Here, we describe six patients with complex rearrangements investigated through a combination of different techniques: karyotyping, chromosomal microarray, and OGM were performed to characterize the rearrangements. Subsequently, lrGS was used to further resolve the alterations, refine their breakpoints' location, and sequence their junction points. Three patients presented CCRs involving three, four, and six chromosomes, while three exhibited CGRs involving one different chromosome each, providing a variety of complex SVs to show the importance of each technique and their combination in rearrangement resolution. In total, the complex rearrangements presented 127 breakpoints, 66 junction points and involved 14 of the 24 chromosomes. Higher-resolution techniques revealed additional complexity in all cases. Despite the advances provided by OGM and lrGS, conventional karyotyping remained indispensable for complete rearrangement resolution. In two patients, the findings supported a novel mechanism combining features of the different chromoanagenesis processes. Furthermore, evidence of inherited alterations was identified, and the comprehensive characterization of the rearrangements enabled more accurate genotype-phenotype correlations. Our findings indicate that an integrated approach combining karyotyping, OGM, and lrGS can completely resolve SVs, including complex rearrangements.
Wang, J.; Chen, J.; Zhao, B.; Zhang, G.; Jian, S.; Deng, T.; Liang, D.
Show abstract
While cycle threshold (Ct) values from quantitative PCR (qPCR) serve as the gold-standard indicators of target abundance, their clinical interpretation is frequently confounded by inherent variability across diverse assay designs, reagents, and instrumentation. In this study, we present a data-driven consensus framework for Ct evaluation that uses large-scale, multi-assay amplification data to establish reference patterns of normal Ct behavior. Based on a total of 41,770 amplification curves collected from four routine diagnostic assays across two PCR platforms, we evaluated machine learning models across three experimental scenarios: within-platform validation, cross-assay generalization, and cross-platform transfer. Extreme gradient boosting (XGBoost) achieved the most accurate and stable predictions under data-sufficient, within-platform conditions with a mean absolute error (MAE) of 0.0419, while pooled multi-assay training improved cross-assay robustness compared with single-assay models. Model performance was further assessed using a deviation-based metric to quantify differences between predicted and instrument-reported Ct values, allowing efficient identification of anomalous amplification curves in large datasets. Notably, direct application across platforms without recalibration led to a substantial decline in performance, with a MAE of 2.62, showing platform-dependent variability. These findings indicate strong stability under within-platform and cross-assay conditions, with scalability contingent upon appropriate cross-platform calibration.
Batista Lozada, Y.; Frometa, Y. G. M.; Gonzalez Gonzalez, Y. J.; Beltran, Y. M.; Garcia de la Rosa, I.; Gutierrez Luis, D.; de Torner, M. L.; Alarcon, A. B.; Triana Mansito, S.; Rodriguez Suarez, A. M.
Show abstract
Background SARS-CoV-2 genomic surveillance is vital for public health, but whole-genome sequencing (WGS) remains costly and inaccessible in many resource-limited settings. We developed and validated a multiplex real-time RT-PCR assay for rapid, economical detection of key mutations associated with variants of interest (VOI) and concern (VOC). Methodology Two multiplex mixes (M1, M2) targeting eight mutations in the ORF1a and Spike genes were designed. Analytical validation included sensitivity, specificity, reproducibility, and limit of detection (LoD) using WHO international standards and a respiratory pathogen panel. In parallel, an in silico analysis evaluated oligonucleotide efficacy against 10.4 million SARS-CoV-2 genomes from GISAID/NCBI, assessing inclusivity, target-site secondary structure (RNAalifold), and hybridization energy (Primer3Plus). Results The assay demonstrated 100% clinical sensitivity among samples with valid RT-PCR results (41/42 samples yielded interpretable results, with one inhibited sample excluded from sensitivity calculation), a LoD of 5.7 log10 IU/mL, and 100% analytical specificity against 32 non-SARS-CoV-2 respiratory pathogens. Six out of eight oligonucleotide sets showed >96% inclusivity; two sets exhibited reduced inclusivity (94.03%, 90.14%) and structural features potentially affecting binding against emerging variants. The assay enables direct identification of major VOCs (Alpha, Beta, Gamma, Delta, Omicron) and indirect detection of multiple VOIs (P.2, Epsilon, Kappa, Eta, Iota, Lambda). Conclusion This standardized multiplex assay provides a rapid, sensitive, and low-cost alternative for SARS-CoV-2 variant surveillance in Cuba and similar settings. The integration of experimental and in silico validation offers a robust, adaptable framework to sustain diagnostic accuracy amid viral evolution, optimizing the allocation of scarce sequencing resources.
Zhang, D.; Wang, Y.; Sager, L.; Koenigsberg, R.; Birari, M.; Hakansson, A.; Fogarty, E.; Reeves, J. W.; Artieri, C.; Lofaro, L.; Russnes, H. G.; Ohnstad, H. O.; Naume, B.; Febbo, P. G.; Marcom, P. K.; Gole, J.
Show abstract
Background: The Prosigna Breast Risk of Recurrence test is based on the PAM50 classifier and was originally validated as an in vitro diagnostic (IVD) test on the Dx enabled nCounter(R) Analysis System. The Prosigna test is intended for early-stage, hormone receptor+ (HR+) breast cancer and provides the risk of recurrence (ROR) score (0-100), intrinsic subtype (Luminal A, Luminal B, HER2-enriched, and Basal-like), and the 10-year probability of distant recurrence. We describe the performance of the Prosigna test as a whole transcriptome RNA sequencing laboratory developed test (LDT) for measuring the Prosigna ROR score and intrinsic subtypes on tissue from surgical resection and core needle biopsy as compared to the Prosigna test on the nCounter system. Methods: We evaluated three separate breast cancer cohorts to 1) bridge the IVD test on the nCounter system and NGS LDT test (n = 245), 2) validate the bridged algorithm on an independent biobank sample set (n = 187), and 3) retrospectively test performance on long-term archival samples from a previous study (n = 109). Results: Bridging analysis showed minimal score variability and robust correlation of Prosigna ROR scoring in surgical resections (SR) (2.459, SD; 0.981, R2) and core needle biopsy (CNB) (2.338, SD; 0.970, R2) samples. In the validation set, the Prosigna NGS LDT ROR scores maintained high correlation to the scores of the nCounter system (SR = 0.968, CNB = 0.966, R2), exhibited minimal score variability (SR = 2.488, CNB = 2.558, SD), and demonstrated high concordance in subtype classifications (SR = 92.3% CNB = 92.8%). Further testing demonstrated comparable performance across tumor fractions, a lower limit of detection (LLOD) of 5 ng, and robustness to exogenous ethanol or genomic DNA contamination. When testing previously extracted RNA from the clinical cohort, we observed high correlation (0.974, R2) and low variance (3.078, SD) of ROR scores with original values on the nCounter system, along with strong risk group (95.4%) and subtype (94.5%) concordance. Conclusions: This study describes the analytical validation of the Prosigna NGS-based LDT measuring the Prosigna ROR score and intrinsic subtypes with robust analytical performance on SR and CNB specimens, providing confidence for clinicians utilizing the NGS-based version of this well-established test.
Allen, S.; Rowlands, C. F.; Kuzbari, Z.; Garrett, A.; Durkie, M.; Burghel, G. J.; Robinson, R.; Callaway, A.; Field, J.; Frugtniet, B.; Palmer-Smith, S.; Grant, J.; Pagan, J.; Johnston, E.; McDevitt, T.; Hughes, L.; Yarram-Smith, L.; Logan, P.; Reed, L.; Snape, K.; McVeigh, T.; Hanson, H.; Roth, F. P.; Starita, L. M.; Fowler, D. M.; Villani, R.; Spurdle, A. B.; Adams, D. J.; Findlay, G.; Turnbull, C.; Cancer Variant Interpretation Group UK (CanVIG-UK),
Show abstract
Background: Clear guidance is lacking regarding how 'truthset' variants should be used for clinical validation of functional assays, namely determining the allocatable evidence points (EPs) towards clinical classification. It is argued that assays should be validated using truthsets of missense variants, as this is the variant type for which classification is most impacted by functional data. EPs will be influenced by both the number of available 'truthset' variants and their concordance with assay readouts. Methods: We first reviewed 112 sets of ClinGen gene-specific classification specifications (CSPECs) to assess methodologies they applied for truthset assembly and clinical validation of assays. We then proposed differing rules regarding variant type and stringency of classification by which truthsets might be assembled using ClinVar-extracted classifications. We then examined augmentation of ClinVar-classified truthsets with 'proxy-clinical' benign-classified missense variants systematically assembled applying ACMG/AMP rules (of differing stringencies). In total, these constituted 70 basic approaches to ClinVar-based truthset assembly, which we applied to VHL, BRCA1, BRCA2 and RAD51C. We additionally analysed the impact on the size of the truthsets of changing the specified phenotypes against which ClinVar classification had been submitted. We then applied these truthsets to quantify concordance and allocatable EPs for five large-scale multiplexed functional assays for VHL, BRCA1, BRCA2, and RAD51C. Results The EPs from clinical validation of each assay varied widely according to which truthset was used across 2,120 permutations of gene-truthset-assay combinations. For example, sequentially applying 700 different ClinVar-based truthsets to 2,268 VHL assay variant readouts (70 basic ClinVar-based approaches, augmented by examining 5 different phenotypes for each basic approach, and separate validation against two defined deleterious zones), the evidence strength allocatable for pathogenicity ranged from nil to strong evidence (0.0 to 5.6 EPs); for benignity it ranged from supporting to strong evidence (-1.5 to -6.4 EPs). Clinical validation using truthsets comprising just ClinVar-classified missense variants typically resulted in lower EPs than truthsets comprising protein truncating (PTV) and synonymous variants; this was more due to paucity of ClinVar-classified missense truthset variants than poorer concordance. Augmentation with larger 'proxy-clinical' benign-classified missense truthsets typically improved evidence allocatable for pathogenicity, with improved power negating modest reduction in concordance. Conclusions EPs can be improved by augmentation with systematically-generated 'proxy-clinical' benign-classified missense variants and/or reduction of truthset stringency. Explicit prescriptive clinical guidance is urgently required to improve consistency in clinical validation of functional assays and consequent evidence application for clinical variant classification.
Yang, Y.; Vasudevaraja, V.; Serrano, J.; Mohamed, H.; Kelly, S.; Jour, G.; Gindin, T.; Park, K.; Jones, D.; Feng, X.; Pinnell, J.; Mclennan, S.; Tin, M. Y.; Tsirigos, A.; Snuderl, M.; Wrzeszczynski, K. O.
Show abstract
Next-generation sequencing (NGS) for the detection of somatic variants has become the method of choice in a variety of molecular oncology fields and in the clinic. Its use ranges from sequencing entire tumor genomes and transcriptomes to targeted clinical diagnostic gene panels. The NYU Langone Genome PACT (Profiling of Actionable Cancer Targets, LG-PACT) assay is a qualitative in vitro diagnostic test that uses targeted next generation sequencing (NGS) of formalin-fixed paraffin-embedded (FFPE) tumor tissue matched with normal specimens from patients to detect gene alterations in a targeted panel covering 606 genes and the TERT promoter. Indications for testing are cancer (solid tumors and hematological malignancies) where a mutational profile from multiple genes would be informative for disease stratification, prognosis, or treatment options including targeted therapies and eligibility for clinical trials. The test is intended to provide information on somatic mutations including point mutations, small insertions/deletions (indels), and copy number aberrations for diagnostic and treatment decisions. LG-PACT is a United States Food and Drug Administration (FDA) cleared diagnostic test (510K: K202304). The clinical interpretation of sequencing data of molecular tumor markers from NGS encompasses automated variant calling tools with human interpretation. This final mostly manual review of data step is intensive, involving highly trained scientists, encompassing literature review, interpretation and clinical tier classification by pathologists, who then provide a complete molecular diagnostic report to the treating oncologists. We provide analysis of 1339 clinical genomic profiles from 31 different cancers and their subtypes, comprising of central nervous system (CNS) 792 (59%) cases (incl. meningioma, glioma and glioblastoma), with 267 (20%) cases predominantly of lung, pancreatic and colorectal and 280 of others (21%). Here, we present the technical challenges of validating an NGS oncological diagnostic targeted assay for clinical grade accuracy and sensitivity for patient care. We show how copy number alterations provide a more comprehensive description of the tumors genomic profile. We then outline the utility of targeted panel sequencing based on certified pathologist selection of reportable variants for our current patient cohort. Where analysis of variant detection has led to 49.4% (661/1339) of our clinical tumor samples containing mutations in known therapy targeted genes, 35.6% (477/1339) with mutation detected in other genes, and 15% (201/1339) cases being negative.
Solanky, D.; Low, C.; Hathaway, C. L.; Cherne, S.; Brown, E.; Palanee-Phillips, T.; Barnabas, R. V.; Bhattacharyya, R. P.; Berdy, B.; Livny, J.
Show abstract
Access to accurate cost-effective technologies for typing high-risk human papillomaviruses (hrHPV) is critical to expand cervical cancer screening and inform vaccination strategies. Compared with clinical-standard quantitative polymerase chain reaction (qPCR) assays, HPV genotyping by next-generation sequencing (NGS) provides greater flexibility, scalability, and genotype specificity. We have developed a method for HPV genotyping, HPV Phased Amplicon Multiplex Sequencing (PhAM-Seq), that uses combinatorial barcoding of amplicons with short, variable-length inline sequences to enable higher throughput and lower per-sample costs than conventional amplicon sequencing approaches. We evaluated HPV PhAM-Seq using degenerate and type-specific primers targeting the L1 and E6-E7 gene loci in a blinded cohort of 170 cervical samples previously typed by the Seegene Anyplex II HPV28 Detection qPCR assay. Across eight common hrHPV types (HPV16, 18, 31, 33, 35, 45, 52, and 58), HPV PhAM-Seq demonstrated >80% overall agreement with qPCR using degenerate L1-targeting primers, with the highest sensitivity for HPV16, 31, 33, and 58. Sensitivity for HPV35, 45, and 52 improved to 85% or greater with type-specific primers targeting genes E6/E7. Parallel processing and sequencing enables a single technician to assay hundreds of samples per week at a reagent cost of around $10 USD per sample, with laboratory automation and sequencing on higher-output platforms enabling further scaling and cost reduction to a scale amenable to population-level surveillance. We include a detailed SOP; tools for primer design, sequencing library construction, and sample tracking; and all scripts needed for data analysis to ensure HPV PhAM-Seq can be readily implemented for scalable, cost-effective hrHPV genotyping or extended to other similar applications.
Maciaszek, J. L.; Pastor Loyola, V.; Cain, T.; Cardenas, M.; Blackburn, P. R.; Wilkinson, M. R.; Koo, S. C.; Wu, C.-H.; Li, C.; Wang, L.; Nichols, K. E.; Klco, J. M.; Eldomery, M. K.
Show abstract
Purpose: Pathogenic or likely pathogenic (P/LP) variants are increasingly identified in genes more commonly associated with adult-onset cancer predisposition, but their prevalence and relevance to children who present with cancer remain unclear. Methods: We retrospectively analyzed 1,280 consecutive pediatric patients with cancer who underwent clinical germline sequencing, using a virtual panel, from 2021 to 2024. Genes with P/LP variants were categorized as aoCPG or pediatric-onset cancer predisposition genes (poCPG) according to cancer risk before age 18 years and pediatric surveillance recommendations. Variant relevance was adjudicated using tumor diagnosis/histopathology, immunohistochemistry, and tumor molecular features and classified as primary, secondary, or indeterminate. Results: Among 1,280 patients, 197 (15.4%) harbored 211 P/LP variants across 54 genes. Sixty-six variants (31.3%) occurred in aoCPG, 87 (41.2%) in poCPG, and 58 (27.5%) were heterozygous variants in autosomal recessive genes. Among adult-onset variants, 7 (10.6%) were primary, 54 (81.8%) secondary, and 5 (7.6%) indeterminate. Among pediatric-onset variants, 77 (88.5%) were primary and 10 (11.5%) secondary. Six patients (3 adult-onset variants; 3 pediatric-onset variants) received targeted therapy informed by germline/somatic sequencing results. Conclusion: In pediatric oncology, most variants in aoCPG are secondary rather than tumor-related findings. Tumor-informed interpretation, beyond variant classification, may improve reporting, counseling, and therapeutic decision-making
Liao, J.; Su, Y.; jiang, F.
Show abstract
Background Current molecular diagnostics for Mycobacterium tuberculosis (MTB) require complex nucleic acid extraction procedures and laboratory infrastructure, limiting their use as point-of-care (POC) tests in resource-limited settings. To address these barriers, we developed an extraction-free one-pot CRISPR assay for rapid detection of MTB directly from minimally processed sputum specimens. Methods The assay integrates ambient-temperature chemical lysis, recombinase polymerase amplification, and CRISPR-Cas12a detection within a single closed-tube workflow. A conserved region of the MTB-specific IS6110 insertion sequence was targeted for detection. Analytical performance was evaluated using serially diluted MTB genomic DNA standards, followed by clinical validation using 100 archived sputum specimens, including 50 MTB-positive samples and 50 MTB-negative controls. Results The extraction-free assay detected MTB genomic DNA within 30 minutes and achieved an analytical limit of detection of 100 copies per reaction. In clinical validation, the assay correctly identified 48 of 50 MTB-positive specimens and 49 of 50 MTB-negative specimens, yielding 96.0% sensitivity and 98.0% specificity. Conclusions This study demonstrates the feasibility of extraction-free one-pot CRISPR-Cas12a detection of MTB directly from sputum specimens. By eliminating conventional nucleic acid extraction while maintaining high analytical sensitivity and diagnostic performance, the platform may facilitate future development of rapid molecular diagnostics for POC and resource-limited settings.
James, L. G.; Thorn, G. J.; Morel, C.; PCRFTB, ; Kocher, H. M.; Ross-Adams, H. E.; Chelala, C.
Show abstract
Whole exome sequencing (WES) of circulating tumour DNA (ctDNA) enables longitudinal monitoring of tumour dynamics, evolution and treatment response but remains technically challenging in low-input, low-shedding settings such as pancreatic ductal adenocarcinoma (PDAC). Here, we systematically compared three commercially available low-input WES workflows incorporating Agilent (V6, V8) and Qiagen exome capture designs using ultra-low input cfDNAs extracted from multiple matched longitudinal plasma samples from PDAC patients. Using predefined performance metrics including coverage, duplication rate and variant detection and additional metrics relevant for clinical genomic profiling in patient care, we show that all three workflows produced high-quality sequencing data, even from very low input cfDNA. Within the conditions tested here, the Agilent V8 workflow provided the most favourable balance of coverage uniformity, sequencing efficiency and hotspot coverage for low input, low tumour fraction cfDNA WES. These findings demonstrate that workflow design, including capture footprint, substantially influences ctDNA WES performance in low-input clinical contexts. These findings are particularly relevant in early stage and/or minimal residual disease settings, where tumour fractions are low and recovery of genomic information from limited-input samples is critical.
Carlomagno, M.; Suarez Lopez, F. J.; Maestri, S.; Esposito, A.; Obadovic, V.; Visconti, V. V.; Ciabini, D.; Marcolungo, L.; Rossi, N.; Casagrande, M.; Angheben, L.; Spadoni, L.; D Apice, M. R.; Novelli, G.; Delledonne, M.; Botta, A.; Rossato, M.
Show abstract
The broader application of long-read sequencing (LRS) for repeat expansion characterization in myotonic dystrophy type 2 (DM2) and other repeat expansion disorders (REDs) remains limited by the lack of systematic validation and benchmarking of sequencing results and bioinformatic workflows. Here, we performed an orthogonal cross-platform validation of previously generated Oxford Nanopore Technologies (ONT) data by sequencing the same DNA samples with Pacific Biosciences (PacBio) HiFi following amplification-free targeted enrichment in a cohort of 8 DM2 patients. Despite substantial differences in sequencing chemistry and coverage, the two platforms showed high concordance in repeat size estimation, somatic mosaicism, and repeat architecture. This validation confirmed the presence of the (TCTG)n motif and enabled the identification of a previously unreported (CCCG)n motif at the 3' end of expanded alleles, further highlighting the structural complexity of the CNBP expansion. Through this analysis, we also established a bioinformatic workflow that improved ONT-based repeat characterization, addressing limitations in motif resolution and enabling more accurate analysis of CNBP expansions. Overall, this study provides a validated framework for LRS-based CNBP repeat analysis, supporting the integration of these technologies into routine molecular investigation for DM2 and other REDs.
Lüth, T.; Schaake, S.; Much, C.; Belyea, M. M.; Seibler, P.; Grünewald, A.; May, P.; Klein, C.; Weissensteiner, H.; Trinh, J.
Show abstract
Background: Low-frequency heteroplasmic mitochondrial DNA (mtDNA) variants are associated with aging and neurological diseases, including Parkinson's disease (PD). Targeted deep mtDNA sequencing using PacBio HiFi long reads has the potential to resolve heteroplasmy across the full mitochondrial genome with high accuracy. Methods: To validate Vega PacBio sequencing for detecting mtDNA heteroplasmy, we analyzed four predefined mixtures of two mtDNA haplotypes. We generated a single long-range PCR amplicon covering the entire mitochondrial genome. These amplicons were mixed at predefined ratios (minor mixture haplotype component: 5%, 2%, 1%, and 0.1%). Variant calling was performed using Mutserve2, and accuracy was assessed by calculating the F1 score from comparisons between expected and detected variants. Full-length mtDNA PacBio sequencing was applied to investigate heteroplasmy across fibroblast passages derived from five LRRK2 p.Gly2019Ser variant carriers (n=3 affected with PD and n=2 unaffected carriers). Changes in mtDNA heteroplasmy level and variant load were assessed longitudinally using a linear mixed model. Results: The single-amplicon approach enabled full-length haplotype resolution without amplification bias associated with overlapping PCR strategies. The F1 score of the predefined mixtures was 1.0 for heteroplasmy levels between 5% and 1% and remained high (0.91) at 0.1%. We detected n=10/62 variants discordant with the Illumina reference at the 0.1% mixture, but sensitivity remained very high at 1.00 in that mixture. Detected minor variants closely matched expected heteroplasmy levels, with average variant levels of 0.057 (5%), 0.022 (2%), 0.011 (1%), and 0.001 (0.1%). Across twelve fibroblast passages, we observed fewer mtDNA heteroplasmic variants ({beta}=-3.2, p=0.026). Increased heteroplasmic variant load over time was also associated with older age ({beta}=1.50, p=0.001) and PD affection status ({beta}=5.0, p=1.0 x 10-4) in LRRK2 variant carriers. Notably, we observed distinct patterns of heteroplasmic variants that either increased or decreased in heteroplasmy level across passages. Conclusion: PacBio HiFi sequencing, combined with a single-amplicon strategy, enables accurate full-length mtDNA heteroplasmy detection and longitudinal analysis, providing a valuable tool for studying mitochondrial variation and dynamics in disease.
Ayati, A.; Onal, G.; Sur, A.; Azzam, S.; Wang, B.; Rudrapatna, V. A.
Show abstract
Objective: Erythropoietic protoporphyria (EPP) is a rare photodermatosis marked by multi-year diagnostic delays. We developed and externally validated machine learning models to identify patients with EPP earlier from longitudinal electronic health record (EHR) data and estimate undiagnosed disease burden. Materials and Methods: In a retrospective case-control study at two San Francisco health systems, an academic referral center (UCSF) and a safety-net hospital (ZSFG) we identified 74 confirmed EPP cases using combined diagnostic coding, biochemical criteria, and specialty chart review. Symptom-enriched controls were sampled at a 40:1 ratio. Longitudinal diagnoses, laboratory results, medications, procedures, and encounters preceding the outcome date were modeled with a gradient-boosting classifier (CatBoost) and a state-space sequence model (MAMBA). The best model was deployed across the UCSF population and externally validated at ZSFG without retraining. Results: On the UCSF held-out test set (n=1,865; 43 cases), MAMBA outperformed CatBoost (AUC ROC 0.91 vs 0.89; average precision 0.42 vs 0.27; precision 65% vs 20%), flagging cases a median of 229 days before documented diagnosis. Deployed across 297,967 symptom-compatible patients, it identified 310 high-risk individuals, implying a prevalence approaching genetic estimates. External validation at ZSFG showed attenuated performance (AUC ROC 0.72; average precision 0.10) while preserving early detection (median 264 days). Discussion: A sequence model integrating temporal EHR signals detected EPP months before clinical recognition, corroborating genetic evidence of substantial underdiagnosis. Cross-site attenuation reflects population and documentation differences and underscores the need for local recalibration. Conclusion: Longitudinal EHR-based machine learning can shorten EPP diagnostic delay and prioritize patients for confirmatory testing, supporting proactive rare-disease case finding.
Hanzlikova, Z.; Styk, J.; Pös, O.; Biro, O.; Bokorova, S.; Lukyova, L.; Sitarcik, J.; Sladecek, T.; Krampl, W.; Meszaros, A.; Hunyadi, P.; Mate, S.; Egeto, A.; Rigo, J.; Sedlackova, T.; Radvanszky, J.; Budis, J.; Szemes, T.
Show abstract
Background: Despite advances in circulating tumor DNA analysis, reliable detection of oncological disease from ultra-low coverage whole genome sequencing (ulcWGS) remains challenging, particularly at low tumor fractions. This study leverages cell-free DNA (cfDNA) characteristics to develop and evaluate a robust, integrative binary predictive model for ovarian cancer (OC) status screening. OC represents a growing global burden and is often diagnosed at advanced stages due to the lack of specific early symptoms and effective screening strategies, highlighting the need for sensitive and broadly applicable early detection approaches.Methods and Findings: We analyzed plasma cfDNA from OC patients (N = 85) and cancer-free controls (N = 41) using ulcWGS (~1x). Participation in the study was voluntary, and all participants provided written informed consent before any study-related procedures under study approval No. 16119-8/2022/EUIG. Within an integrated workflow combining standardized laboratory processing, bioinformatic pipelines, and machine learning (ML), we extracted 21 features capturing copy number variations (CNVs) and fragmentomic characteristics to identify complementary signatures distinguishing OC from controls. Predictive models were developed using XGBoost with hyperparameter optimization and evaluated on an independent test set (n = 25% of the cohort). A dual-threshold classification strategy was applied to define an uncertainty zone and optimize screening performance.CNV-derived and fragmentomic features assessed in exploratory analysis on the training-validation set showed moderate discriminative power (AUC 0.569 - 0.946) but substantial overlap between groups. On the test set, the CNV-only model achieved an AUC of 0.855 (sensitivity 85%, specificity 50%), while the fragmentomics-only model reached an AUC of 0.8825 (sensitivity 95%, specificity 30%). Both feature domains captured complementary aspects of tumor-derived cfDNA, with fragmentomics favoring sensitivity and CNV-derived metrics improving specificity. Integration of both feature classes improved performance, yielding an AUC of 0.900, sensitivity of 85.00%, and specificity of 90.00%. SHAP analysis confirmed contributions from both feature types without a single dominant predictor.Conclusions: We present an integrative cfDNA framework for OC detection based on ulcWGS that combines CNV and fragmentomic signals to improve diagnostic performance over single-feature approaches. By enabling robust detection of tumor-associated patterns at ultra-low sequencing depth, this approach demonstrates that meaningful cancer discrimination can be achieved without reliance on deep sequencing. This highlights the potential of cost-effective and scalable liquid biopsy strategies for population-level cancer screening their integration into personalized and preventive oncology. Keywords: Liquid biopsy, ovarian cancer, ultra-low coverage whole genome sequencing, cell-free DNA, cell-free tumor DNA, cancer detection, copy number variations, insert size, fragmentomics, machine learning
Allen, S.; Rowlands, C. F.; Garrett, A.; Kuzbari, Z.; Durkie, M.; Burghel, G. J.; Robinson, R.; Callaway, A.; Field, J.; Frugtniet, B.; Palmer-Smith, S.; Grant, J.; Pagan, J.; Johnston, E.; McDevitt, T.; Hughes, L.; Yarram-Smith, L.; Logan, P.; Reed, L.; Snape, K.; McVeigh, T.; Hanson, H.; Villani, R.; Spurdle, A. B.; Starita, L. M.; Fowler, D. M.; Roth, F. P.; Radford, E.; Adams, D. J.; Findlay, G. M.; Turnbull, C.; Cancer Variant Interpretation Group UK (CanVIG-UK),
Show abstract
Background Large-scale functional assays, including multiplex assays of variant effect, have substantial potential to resolve variants of uncertain significance (VUS), particularly for rare missense variants where clinical and population evidence are limited. The ClinGen assay-level clinical validation framework described by Brnich et al provided baseline guidance for the use of functional data for variant classification. However, clear consensus regarding construction of variant 'truthsets' by which to clinically validate functional data remains lacking. Methods CanVIG-UK developed consensus recommendations for truthset construction through an iterative national consultation process involving the CanVIG Steering Advisory Group (CStAG), wider CanVIG-UK membership, and engagement with international functional genomics experts. Consultation was based on previous analyses of 2,120 truthset constructions examining the impact of truthset composition on evidence point allocation within the ClinGen assay-level clinical validation framework. Results Across several consultations, CanVIG-UK established nine guiding principles and seven best-practice recommendations for assay-level clinical validation, using the assumed context of an assay for a cancer susceptibility gene where loss-of-function is the mechanism of pathogenicity. The principal recommendation stipulates, where assays are intended for use in interpretation of largely missense variants, the truthset used to validate should comprise only missense variants. Rather than mixtures of different variant types which may serve to over-estimate assay performance. Additional recommendations support option for relaxation of truthset stringency to improve power, augmentation of benign missense truthsets with systematically derived 'proxy-clinical' benign variants, independent clinical validation separate from assayist-defined validation, and careful evaluation of missense score distributions against that of protein-truncating and synonymous variants. Guidance is also provided for scenarios with limited pathogenic truthset availability and for assays reporting multiple deleterious zones or readouts. Conclusions The CanVIG-UK principles and recommendations for truthset construction upon the ClinGen assay-level clinical validation framework, while aiming to form a baseline for future discussion regarding other functional and disease contexts and helping to address the gap between publication of new data and routine clinical implementation.
Mabvakure, B. M.; Promprasert, P.; Martinez Cruz, L.; Patil, S.; Barros, J.; Hayhurst, M.; Mohebbi, E.; de la Caridad Delgado Herrera, D.; Lee, G. J.; Latif, S.; Williams, F.; Samdani, R.; Duttargi, A.; Berhane, B.; Besufikad, E.; Tadesse, S.; Jibril Suleiman, A.; Lefante, C.; Hsieh, M.-C.; Purrington, K.; Adjei, E.; Qin, T.; Sartor, M.; Stoffel, E. M.; Rozek, L. S.
Show abstract
PURPOSE Colorectal cancer (CRC) incidence and mortality rates differ by population, and evidence suggests that genetic differences may affect cancer biology. However, studies investigating CRC variants in genetically heterogeneous populations are limited. Using somatic tumor mutation profiling of CRCs diagnosed in African Americans (AAs), Ghanaians, Ethiopians, and NHWs, we explore correlations between population group and population-specific tumor variants. PATIENTS AND METHODS Somatic DNA from CRC tumors resected from 150 individuals, including 43 AAs (27%), 53 NHWs (35%), 21 Ghanaians (14.2%), and 33 Ethiopians (22.3%), was sequenced on the Illumina NovaSeq platform, targeting 290 genes. We compared mutations in AAs, Ghanaians, and Ethiopians to those in NHWs to identify variants enriched in historically underrepresented groups. RESULTS US cohort tumors were diagnosed at significantly younger ages with more early-onset cases (<50 years old) than African cohorts (p <0.05). Significant differences were observed in primary tumor location, MMR phenotypes, KRAS mutations, and distribution of tumor mutational burden by population. BRAF V600E mutations were rare across all groups, while non-V600E BRAF mutation rates were higher in AA and NHW (43-44%) than Ethiopian and Ghanaian (14-33%) samples. Population-specific differences were identified in mutation rates of APC, CTNNB1, RNF43, PIK3CA, and TP53, as well as in pathogenic variant occurrence.
Ahmed, A. F. F.
Show abstract
Background Lung cancer mortality is rising in Libya, but access to molecular diagnostics for EGFR mutations--essential for guiding tyrosine kinase inhibitor therapy--remains severely limited. Selecting an appropriate testing platform requires balancing analytical performance against cost and infrastructure constraints. Methods We conducted a prospective comparative validation study using formalin-fixed paraffin-embedded (FFPE) tissue samples from Libyan non-small cell lung cancer (NSCLC) patients. Following stringent DNA quality control, samples were tested in parallel across four platforms: multiplex real-time PCR (MRT-PCR), reverse hybridization strip assay (RHSA), agarose gel electrophoresis (AGE), and immunohistochemistry (IHC). Performance was assessed by inter-method concordance, turnaround time, and cost per test. Results Of 30 initial samples, only six (20%) met quality thresholds (A260/A280 1.70-1.90; concentration [≥]10 ng/{micro}L), highlighting pre-analytical challenges. Three samples harbored EGFR exon 19 deletions. A critical discordance was identified: one sample tested negative by MRT-PCR (Ct {approx}38, {Delta}Ct=13) but positive by RHSA, AGE, and IHC, indicating a false-negative result from the reference method. IHC and RHSA offered the most favorable balance of cost (USD 40-75/test) and operational feasibility, while MRT-PCR (USD 150/test) required specialized infrastructure. Conclusions Relying solely on automated PCR may lead to under-diagnosis in low-cellularity or degraded FFPE samples. We recommend a hybrid algorithm: IHC as a cost-effective primary screen, followed by RHSA for confirmation. This approach optimizes resource allocation and improves diagnostic equity in Libya.
Vargas-Reyes, M.; Alcantara, R.; Herrera, C.; Townsend, M.; Flores-Jimenes, K.; Raymundo, C.; Milon, P.
Show abstract
Antimicrobial resistance (AMR) represents a major global health threat, with plasmid-borne mcr genes driving colistin resistance and exposing critical gaps in One-Health surveillance across human, animal, and environmental reservoirs. The most prevalent variant, mcr-1, remains difficult to monitor in resource-limited settings due to the lack of rapid, affordable, and field-deployable molecular tools. Here, we developed C12amcr, an integrated molecular toolbox that combines pre-amplification PCR with a fluorescent CRISPR-Cas12a assay targeting a conserved region of mcr-1 and a custom low-cost, hand-held 3D-printed portable fluorometer. Under optimized conditions, the assay achieved a limit of detection of 630 cells/mL. In poultry feces spiked with mcr-1-positive E. coli, C12amcr detected as few as 1,800 cells/mL. When tested on 22 community-derived E. coli isolates, the assay showed 100% concordance with both next-generation sequencing for mcr-1 detection and phenotypic colistin susceptibility testing by broth microdilution. The accompanying portable fluorometer performed equivalently to a laboratory microplate reader while enabling fully decentralized workflows compatible with portable PCR platforms. By integrating locally produced molecular reagents, straightforward protocols, and an accessible field-ready fluorescence reader, C12amcr overcomes key barriers to decentralized AMR surveillance and provides a practical, scalable solution for One-Health monitoring in resource-limited settings.